176 research outputs found

    Survey of Vector Database Management Systems

    Full text link
    There are now over 20 commercial vector database management systems (VDBMSs), all produced within the past five years. But embedding-based retrieval has been studied for over ten years, and similarity search a staggering half century and more. Driving this shift from algorithms to systems are new data intensive applications, notably large language models, that demand vast stores of unstructured data coupled with reliable, secure, fast, and scalable query processing capability. A variety of new data management techniques now exist for addressing these needs, however there is no comprehensive survey to thoroughly review these techniques and systems. We start by identifying five main obstacles to vector data management, namely vagueness of semantic similarity, large size of vectors, high cost of similarity comparison, lack of natural partitioning that can be used for indexing, and difficulty of efficiently answering hybrid queries that require both attributes and vectors. Overcoming these obstacles has led to new approaches to query processing, storage and indexing, and query optimization and execution. For query processing, a variety of similarity scores and query types are now well understood; for storage and indexing, techniques include vector compression, namely quantization, and partitioning based on randomization, learning partitioning, and navigable partitioning; for query optimization and execution, we describe new operators for hybrid queries, as well as techniques for plan enumeration, plan selection, and hardware accelerated execution. These techniques lead to a variety of VDBMSs across a spectrum of design and runtime characteristics, including native systems specialized for vectors and extended systems that incorporate vector capabilities into existing systems. We then discuss benchmarks, and finally we outline research challenges and point the direction for future work.Comment: 25 page

    Multi-ancestry genome-wide study in >2.5 million individuals reveals heterogeneity in mechanistic pathways of type 2 diabetes and complications

    Get PDF
    Type 2 diabetes (T2D) is a heterogeneous disease that develops through diverse pathophysiological processes. To characterise the genetic contribution to these processes across ancestry groups, we aggregate genome-wide association study (GWAS) data from 2,535,601 individuals (39.7% non-European ancestry), including 428,452 T2D cases. We identify 1,289 independent association signals at genome-wide significance (P&lt;5×10 - 8 ) that map to 611 loci, of which 145 loci are previously unreported. We define eight non-overlapping clusters of T2D signals characterised by distinct profiles of cardiometabolic trait associations. These clusters are differentially enriched for cell-type specific regions of open chromatin, including pancreatic islets, adipocytes, endothelial, and enteroendocrine cells. We build cluster-specific partitioned genetic risk scores (GRS) in an additional 137,559 individuals of diverse ancestry, including 10,159 T2D cases, and test their association with T2D-related vascular outcomes. Cluster-specific partitioned GRS are more strongly associated with coronary artery disease and end-stage diabetic nephropathy than an overall T2D GRS across ancestry groups, highlighting the importance of obesity-related processes in the development of vascular outcomes. Our findings demonstrate the value of integrating multi-ancestry GWAS with single-cell epigenomics to disentangle the aetiological heterogeneity driving the development and progression of T2D, which may offer a route to optimise global access to genetically-informed diabetes care. </p

    Association analyses of East Asian individuals and trans-ancestry analyses with European individuals reveal new loci associated with cholesterol and triglyceride levels

    Get PDF
    Large-scale meta-analyses of genome-wide association studies (GWAS) have identified >175 loci associated with fasting cholesterol levels, including total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), and triglycerides (TG). With differences in linkage disequilibrium (LD) structure and allele frequencies between ancestry groups, studies in additional large samples may detect new associations. We conducted staged GWAS meta-analyses in up to 69,414 East Asian individuals from 24 studies with participants from Japan, the Philippines, Korea, China, Singapore, and Taiwan. These meta-analyses identified (P < 5 × 10-8) three novel loci associated with HDL-C near CD163-APOBEC1 (P = 7.4 × 10-9), NCOA2 (P = 1.6 × 10-8), and NID2-PTGDR (P = 4.2 × 10-8), and one novel locus associated with TG near WDR11-FGFR2 (P = 2.7 × 10-10). Conditional analyses identified a second signal near CD163-APOBEC1. We then combined results from the East Asian meta-analysis with association results from up to 187,365 European individuals from the Global Lipids Genetics Consortium in a trans-ancestry meta-analysis. This analysis identified (log10Bayes Factor ≥6.1) eight additional novel lipid loci. Among the twelve total loci identified, the index variants at eight loci have demonstrated at least nominal significance with other metabolic traits in prior studies, and two loci exhibited coincident eQTLs (P < 1 × 10-5) in subcutaneous adipose tissue for BPTF and PDGFC. Taken together, these analyses identified multiple novel lipid loci, providing new potential therapeutic targets

    Universal screening for Lynch syndrome in a large consecutive cohort of Chinese colorectal cancer patients: High prevalence and unique molecular features

    Get PDF
    The prevalence of Lynch syndrome (LS) varies significantly in different populations, suggesting that ethnic features might play an important role. We enrolled 3330 consecutive Chinese patients who had surgical resection for newly diagnosed colorectal cancer. Universal screening for LS was implemented, including immunohistochemistry for mismatch repair (MMR) proteins, BRAFV600E mutation test and germline sequencing. Among the 3250 eligible patients, MMR protein deficiency (dMMR) was detected in 330 (10.2%) patients. Ninety‐three patients (2.9%) were diagnosed with LS. Nine (9.7%) patients with LS fulfilled Amsterdam criteria II and 76 (81.7%) met the revised Bethesda guidelines. Only 15 (9.7%) patients with absence of MLH1 on IHC had BRAFV600E mutation. One third (33/99) of the MMR gene mutations have not been reported previously. The age of onset indicates risk of LS in patients with dMMR tumors. For patients older than 65 years, only 2 patients (5.7%) fulfilling revised Bethesda guidelines were diagnosed with LS. Selective sequencing of all cases with dMMR diagnosed at or below age 65 years and only of those dMMR cases older than 65 years who fulfill revised Bethesda guidelines results in 8.2% fewer cases requiring germline testing without missing any LS diagnoses. While the prevalence of LS in Chinese patients is similar to that of Western populations, the spectrum of constitutional mutations and frequency of BRAFV600E mutation is different. Patients older than 65 years who do not meet the revised Bethesda guidelines have a low risk of LS, suggesting germline sequencing might not be necessary in this population

    Finishing the euchromatic sequence of the human genome

    Get PDF
    The sequence of the human genome encodes the genetic instructions for human physiology, as well as rich information about human evolution. In 2001, the International Human Genome Sequencing Consortium reported a draft sequence of the euchromatic portion of the human genome. Since then, the international collaboration has worked to convert this draft into a genome sequence with high accuracy and nearly complete coverage. Here, we report the result of this finishing process. The current genome sequence (Build 35) contains 2.85 billion nucleotides interrupted by only 341 gaps. It covers ∼99% of the euchromatic genome and is accurate to an error rate of ∼1 event per 100,000 bases. Many of the remaining euchromatic gaps are associated with segmental duplications and will require focused work with new methods. The near-complete sequence, the first for a vertebrate, greatly improves the precision of biological analyses of the human genome including studies of gene number, birth and death. Notably, the human enome seems to encode only 20,000-25,000 protein-coding genes. The genome sequence reported here should serve as a firm foundation for biomedical research in the decades ahead

    Guidelines for the use and interpretation of assays for monitoring autophagy (3rd edition)

    Get PDF
    In 2008 we published the first set of guidelines for standardizing research in autophagy. Since then, research on this topic has continued to accelerate, and many new scientists have entered the field. Our knowledge base and relevant new technologies have also been expanding. Accordingly, it is important to update these guidelines for monitoring autophagy in different organisms. Various reviews have described the range of assays that have been used for this purpose. Nevertheless, there continues to be confusion regarding acceptable methods to measure autophagy, especially in multicellular eukaryotes. For example, a key point that needs to be emphasized is that there is a difference between measurements that monitor the numbers or volume of autophagic elements (e.g., autophagosomes or autolysosomes) at any stage of the autophagic process versus those that measure fl ux through the autophagy pathway (i.e., the complete process including the amount and rate of cargo sequestered and degraded). In particular, a block in macroautophagy that results in autophagosome accumulation must be differentiated from stimuli that increase autophagic activity, defi ned as increased autophagy induction coupled with increased delivery to, and degradation within, lysosomes (inmost higher eukaryotes and some protists such as Dictyostelium ) or the vacuole (in plants and fungi). In other words, it is especially important that investigators new to the fi eld understand that the appearance of more autophagosomes does not necessarily equate with more autophagy. In fact, in many cases, autophagosomes accumulate because of a block in trafficking to lysosomes without a concomitant change in autophagosome biogenesis, whereas an increase in autolysosomes may reflect a reduction in degradative activity. It is worth emphasizing here that lysosomal digestion is a stage of autophagy and evaluating its competence is a crucial part of the evaluation of autophagic flux, or complete autophagy. Here, we present a set of guidelines for the selection and interpretation of methods for use by investigators who aim to examine macroautophagy and related processes, as well as for reviewers who need to provide realistic and reasonable critiques of papers that are focused on these processes. These guidelines are not meant to be a formulaic set of rules, because the appropriate assays depend in part on the question being asked and the system being used. In addition, we emphasize that no individual assay is guaranteed to be the most appropriate one in every situation, and we strongly recommend the use of multiple assays to monitor autophagy. Along these lines, because of the potential for pleiotropic effects due to blocking autophagy through genetic manipulation it is imperative to delete or knock down more than one autophagy-related gene. In addition, some individual Atg proteins, or groups of proteins, are involved in other cellular pathways so not all Atg proteins can be used as a specific marker for an autophagic process. In these guidelines, we consider these various methods of assessing autophagy and what information can, or cannot, be obtained from them. Finally, by discussing the merits and limits of particular autophagy assays, we hope to encourage technical innovation in the field

    Large expert-curated database for benchmarking document similarity detection in biomedical literature search

    Get PDF
    Document recommendation systems for locating relevant literature have mostly relied on methods developed a decade ago. This is largely due to the lack of a large offline gold-standard benchmark of relevant documents that cover a variety of research fields such that newly developed literature search techniques can be compared, improved and translated into practice. To overcome this bottleneck, we have established the RElevant LIterature SearcH consortium consisting of more than 1500 scientists from 84 countries, who have collectively annotated the relevance of over 180 000 PubMed-listed articles with regard to their respective seed (input) article/s. The majority of annotations were contributed by highly experienced, original authors of the seed articles. The collected data cover 76% of all unique PubMed Medical Subject Headings descriptors. No systematic biases were observed across different experience levels, research fields or time spent on annotations. More importantly, annotations of the same document pairs contributed by different scientists were highly concordant. We further show that the three representative baseline methods used to generate recommended articles for evaluation (Okapi Best Matching 25, Term Frequency-Inverse Document Frequency and PubMed Related Articles) had similar overall performances. Additionally, we found that these methods each tend to produce distinct collections of recommended articles, suggesting that a hybrid method may be required to completely capture all relevant articles. The established database server located at https://relishdb.ict.griffith.edu.au is freely available for the downloading of annotation data and the blind testing of new methods. We expect that this benchmark will be useful for stimulating the development of new powerful techniques for title and title/abstract-based search engines for relevant articles in biomedical research.Peer reviewe

    Height and body-mass index trajectories of school-aged children and adolescents from 1985 to 2019 in 200 countries and territories: a pooled analysis of 2181 population-based studies with 65 million participants

    Get PDF
    Summary Background Comparable global data on health and nutrition of school-aged children and adolescents are scarce. We aimed to estimate age trajectories and time trends in mean height and mean body-mass index (BMI), which measures weight gain beyond what is expected from height gain, for school-aged children and adolescents. Methods For this pooled analysis, we used a database of cardiometabolic risk factors collated by the Non-Communicable Disease Risk Factor Collaboration. We applied a Bayesian hierarchical model to estimate trends from 1985 to 2019 in mean height and mean BMI in 1-year age groups for ages 5–19 years. The model allowed for non-linear changes over time in mean height and mean BMI and for non-linear changes with age of children and adolescents, including periods of rapid growth during adolescence. Findings We pooled data from 2181 population-based studies, with measurements of height and weight in 65 million participants in 200 countries and territories. In 2019, we estimated a difference of 20 cm or higher in mean height of 19-year-old adolescents between countries with the tallest populations (the Netherlands, Montenegro, Estonia, and Bosnia and Herzegovina for boys; and the Netherlands, Montenegro, Denmark, and Iceland for girls) and those with the shortest populations (Timor-Leste, Laos, Solomon Islands, and Papua New Guinea for boys; and Guatemala, Bangladesh, Nepal, and Timor-Leste for girls). In the same year, the difference between the highest mean BMI (in Pacific island countries, Kuwait, Bahrain, The Bahamas, Chile, the USA, and New Zealand for both boys and girls and in South Africa for girls) and lowest mean BMI (in India, Bangladesh, Timor-Leste, Ethiopia, and Chad for boys and girls; and in Japan and Romania for girls) was approximately 9–10 kg/m2. In some countries, children aged 5 years started with healthier height or BMI than the global median and, in some cases, as healthy as the best performing countries, but they became progressively less healthy compared with their comparators as they grew older by not growing as tall (eg, boys in Austria and Barbados, and girls in Belgium and Puerto Rico) or gaining too much weight for their height (eg, girls and boys in Kuwait, Bahrain, Fiji, Jamaica, and Mexico; and girls in South Africa and New Zealand). In other countries, growing children overtook the height of their comparators (eg, Latvia, Czech Republic, Morocco, and Iran) or curbed their weight gain (eg, Italy, France, and Croatia) in late childhood and adolescence. When changes in both height and BMI were considered, girls in South Korea, Vietnam, Saudi Arabia, Turkey, and some central Asian countries (eg, Armenia and Azerbaijan), and boys in central and western Europe (eg, Portugal, Denmark, Poland, and Montenegro) had the healthiest changes in anthropometric status over the past 3·5 decades because, compared with children and adolescents in other countries, they had a much larger gain in height than they did in BMI. The unhealthiest changes—gaining too little height, too much weight for their height compared with children in other countries, or both—occurred in many countries in sub-Saharan Africa, New Zealand, and the USA for boys and girls; in Malaysia and some Pacific island nations for boys; and in Mexico for girls. Interpretation The height and BMI trajectories over age and time of school-aged children and adolescents are highly variable across countries, which indicates heterogeneous nutritional quality and lifelong health advantages and risks

    Repositioning of the global epicentre of non-optimal cholesterol

    Get PDF
    High blood cholesterol is typically considered a feature of wealthy western countries(1,2). However, dietary and behavioural determinants of blood cholesterol are changing rapidly throughout the world(3) and countries are using lipid-lowering medications at varying rates. These changes can have distinct effects on the levels of high-density lipoprotein (HDL) cholesterol and non-HDL cholesterol, which have different effects on human health(4,5). However, the trends of HDL and non-HDL cholesterol levels over time have not been previously reported in a global analysis. Here we pooled 1,127 population-based studies that measured blood lipids in 102.6 million individuals aged 18 years and older to estimate trends from 1980 to 2018 in mean total, non-HDL and HDL cholesterol levels for 200 countries. Globally, there was little change in total or non-HDL cholesterol from 1980 to 2018. This was a net effect of increases in low- and middle-income countries, especially in east and southeast Asia, and decreases in high-income western countries, especially those in northwestern Europe, and in central and eastern Europe. As a result, countries with the highest level of non-HDL cholesterol-which is a marker of cardiovascular riskchanged from those in western Europe such as Belgium, Finland, Greenland, Iceland, Norway, Sweden, Switzerland and Malta in 1980 to those in Asia and the Pacific, such as Tokelau, Malaysia, The Philippines and Thailand. In 2017, high non-HDL cholesterol was responsible for an estimated 3.9 million (95% credible interval 3.7 million-4.2 million) worldwide deaths, half of which occurred in east, southeast and south Asia. The global repositioning of lipid-related risk, with non-optimal cholesterol shifting from a distinct feature of high-income countries in northwestern Europe, north America and Australasia to one that affects countries in east and southeast Asia and Oceania should motivate the use of population-based policies and personal interventions to improve nutrition and enhance access to treatment throughout the world.Peer reviewe
    corecore